Papers with external pretraining data
Downstream Datasets Make Surprisingly Good Pretraining Corpora (2023.acl-long)
Copied to clipboard
| Challenge: | a dominant practice is to fine tune large pretrained transformer models using smaller downstream datasets . performance gains are not always attributable to the use of external data in massive amounts . |
| Approach: | They propose to use the same (downstream) training data for pretraining and finetuning to compare models. |
| Outcome: | The proposed model outperforms standard pretraining on the BookWiki corpus on 7 and 5 datasets. |